Papers with causal inference

31 papers
Measuring and Modeling Language Change (N19-5)

Copied to clipboard

Challenge: This tutorial will help researchers answer questions fundamental to the social sciences and humanities .
Approach: This tutorial is designed to help researchers answer questions in the social sciences and humanities . it synthesizes recent computational techniques for handling and modeling temporal data .
Outcome: The tutorial will synthesize recent techniques for handling and modeling temporal data, such as dynamic word embeddings, and identify useful tools for social scientists and digital humanities scholars.
Robust Hate Speech Detection via Mitigating Spurious Correlations (2022.aacl-short)

Copied to clipboard

Challenge: a novel hate speech detection model can be used to detect word- and character-level adversarial attacks . existing adversarials assume that attackers replace the target words with other names to evade detection .
Approach: They propose a robust hate speech detection model that can defend against adversarial attacks . they describe the process of hate speech recognition by a causal graph and a regularized entropy loss function to quantify spurious correlation .
Outcome: The proposed model can defend against word- and character-level adversarial attacks.
Causal Investigation of Public Opinion during the COVID-19 Pandemic via Social Media Text (2022.lrec-1)

Copied to clipboard

Challenge: Social distancing orders are the most effective strategy to reduce the spread of COVID-19 in 2020 .
Approach: They propose to use NLP methods in a causal mediation scenario to emphasize the use of NLP and economics to decouple the effect of government restrictions on mobility from the effect that occurs due to public perception of the COVID-19 strategy.
Outcome: The proposed model decouples the effect of government restrictions on mobility behavior from the effect that occurs due to public perception of the COVID-19 strategy in a country.
Mining the Cause of Political Decision-Making from Social Media: A Case Study of COVID-19 Policies across the US States (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on political responsiveness focus on long-term policies collected over decades . recent COVID-19 pandemic has given rise to a new political phenomenon, where political leaders make frequent short-term decisions on the same controlled topic.
Approach: They propose to use Twitter data to classify the sentiments toward governors of each state and conduct controlled studies and comparisons.
Outcome: The proposed model focuses on the COVID-19 pandemic, where political leaders make frequent short-term decisions on the same controlled topic.
Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and Beyond (2022.tacl-1)

Copied to clipboard

Challenge: causality has not had the same importance in natural language processing, says aaron e. smith . he says research on causality in NLP remains scattered across domains without unified definitions .
Approach: They propose to consolidate research on causality in NLP across academic areas . they explore potential uses of causal inference to improve robustness, fairness, interpretability .
Outcome: The proposed method is a unified overview of causal inference for the NLP community.
Text-Transport: Toward Learning Causal Effects of Natural Language (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for causal inference require strong assumptions about the data, meaning the data from which one *can* estimate valid causal effects is not representative of the actual target domain of interest.
Approach: They propose a method for estimation of causal effects from natural language under any text distribution using the notion of distribution shift.
Outcome: The proposed method can be used to estimate causal effects from natural language under any text distribution.
Quantifying the Impact of Structured Output Format on Large Language Models through Causal Inference (2026.findings-eacl)

Copied to clipboard

Challenge: Prior studies have examined the impact of structured output on LLMs’ generation quality, often presenting one-way findings.
Approach: They propose to derive five potential causal structures characterizing the influence of structured output on LLMs’ generation using one assumed and two guaranteed constraints.
Outcome: The proposed pipeline can be extended to other modules and is not limited to structured output but can be used in industrial applications.
XLTime: A Cross-Lingual Knowledge Transfer Framework for Temporal Expression Extraction (2022.findings-naacl)

Copied to clipboard

Challenge: Temporal Expression Extraction (TEE) is essential for understanding time in natural language.
Approach: They propose a framework for multilingual Temporal Expression Extraction that leverages pre-trained language models to prompt cross-language knowledge transfer from English to non-English languages.
Outcome: The proposed framework outperforms the existing SOTA methods on French, Spanish, Portuguese, and Basque by large margins.
Everything Has a Cause: Leveraging Causal Inference in Legal Text Analysis (2021.naacl-main)

Copied to clipboard

Challenge: Existing studies focus on analyzing structured data, while mining causal relationship among factors from unstructured data is of great importance.
Approach: They propose a graph-based causal inference framework which builds causal graphs from fact descriptions without much human involvement.
Outcome: The proposed framework can capture nuance from fact descriptions among confusing charges and provide explainable discrimination in few-shot settings.
CaM-Gen: Causally Aware Metric-Guided Text Generation (2022.findings-acl)

Copied to clipboard

Challenge: Content is created for a well-defined purpose, often described by a metric or signal . external metrics and content tend to have inherent relationships and not all of them may be of consequence.
Approach: They propose a mechanism to guide generative models by user-defined target metrics . authors propose generative networks guided by causally significant aspects of text .
Outcome: The proposed models beat baselines in terms of the target metric control while maintaining fluency and language quality of the generated text.
Exploring Logically Dependent Multi-task Learning with Causal Inference (2020.emnlp-main)

Copied to clipboard

Challenge: Hierarchical multi-task learning models can utilize task dependencies by stacking encoders and outperform democratic ones.
Approach: They propose a model that utilizes the labels of all lower-level tasks and a Gumbel sampling model to deal with cascading errors.
Outcome: The proposed model outperforms democratic models on six out of seven subtasks and achieves state-of-the-art on the two English and one Chinese datasets.
Looking at Radiology Report Generation through a Causal Lens: A Survey (2026.acl-long)

Copied to clipboard

Challenge: Existing surveys on RRG emphasize deep learning while overlooking the critical role of causality.
Approach: They propose to analyze biases across the RRG pipeline and formalize it as a causal modeling problem and review representative causal techniques from the literature.
Outcome: The proposed model can mitigate biases and yield fair, reliable systems with clinically meaningful outputs.
Distilling Causal Effect from Miscellaneous Other-Class for Continual Named Entity Recognition (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for Named Entity Recognition (NER) are not able to learn Other-Class in the same way as new entity types.
Approach: They propose a unified causal framework to retrieve causality from new entity types and Other-Class.
Outcome: The proposed method outperforms the state-of-the-art method on three benchmark datasets.
Multimodal Causal Reasoning Benchmark: Challenging Multimodal Large Language Models to Discern Causal Links Across Modalities (2025.findings-acl)

Copied to clipboard

Challenge: Existing MLLMs lack robustness in multimodal causal reasoning compared to their performance in textual settings.
Approach: They propose a novel multimodal chain-of-thought (CoT) reasoning benchmark that leverages siamese images and text pairs to challenge MLLMs.
Outcome: The proposed benchmark leverages siamese images and text pairs to challenge MLLMs.
Causal Inference with Large Language Model: A Survey (2025.findings-naacl)

Copied to clipboard

Challenge: Existing causal inference frameworks do not match human judgment in several key areas, such as domain knowledge, logical inference, and cultural context.
Approach: They propose to apply large language models to causal inference tasks . they summarize the main causal problems and approaches and compare their results .
Outcome: The proposed methods are compared with traditional methods in healthcare, finance, and economics.
CCG: Rare-Label Prediction via Neural SEM–Driven Causal Game (2025.findings-emnlp)

Copied to clipboard

Challenge: Multi-label classification (MLC) faces persistent challenges from label imbalance, spurious correlations, distribution shifts, especially in rare label prediction.
Approach: They propose a Causal Cooperative Game framework that models multi-player cooperative process for multi-label classification.
Outcome: The proposed framework improves rare label prediction and overall robustness compared to baselines.
Causal Denoising Prototypical Network for Few-Shot Multi-label Aspect Category Detection (2025.findings-acl)

Copied to clipboard

Challenge: Recent methods that learn robust prototypes to represent aspects with limited support samples address noise categories in the support set that hinder their models from effective prototype generation.
Approach: They propose a causal denoising prototypical network for few-shot MACD by learning robust prototypes to represent categories with limited support samples.
Outcome: The proposed model outperforms baseline models and can prevent models from overly predicting more categories and mitigate semantic ambiguity issues among categories.
Distilling Causal Effect of Data in Continual Few-shot Relation Learning (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for learning relational patterns from data are prone to catastrophic forgetting issues due to limited number of samples and continual training mode.
Approach: They propose a unified causal framework for CFRL to restore causal effects from old data . they establish two additional causal paths from old to predictions by colliding with old data separately in the old feature space.
Outcome: The proposed method is superior to existing state-of-the-art methods in CFRL task settings.
Challenges of Using Text Classifiers for Causal Inference (D18-1)

Copied to clipboard

Challenge: a number of scientific analyses focus on low-dimensional structured data, but text classifiers can be used to produce structured variables.
Approach: They propose to use text classifiers to conduct causal analyses on simulated and Yelp data.
Outcome: The proposed method can be used on simulated and Yelp data.
Enhancing Factual Consistency in Text Summarization via Counterfactual Debiasing (2025.coling-main)

Copied to clipboard

Challenge: Abstractive text summarization has produced fluent and informative outputs, but factual inconsistency is a challenge.
Approach: They propose a framework that mitigates the causal effects of language bias and irrelevancy bias by counterfactual estimation.
Outcome: The proposed framework outperforms baseline methods on two widely used summarization datasets.
Enhancing Zero-Shot Chain-of-Thought Reasoning in Large Language Models through Logic (2024.lrec-main)

Copied to clipboard

Challenge: Experimental evaluations of large language models demonstrate the efficacy of enhanced reasoning by logic.
Approach: They propose a framework that uses symbolic logic to verify and rectify reasoning steps by steps.
Outcome: The proposed framework improves the zero-shot chain-of-thought reasoning ability of large language models by verifying and rectifying the reasoning steps step by step.
Multimodal Clickbait Detection by De-confounding Biases Using Causal Representation Inference (2024.emnlp-main)

Copied to clipboard

Challenge: a new method to detect clickbait posts on the Web is needed to detect such posts.
Approach: They propose a method to detect clickbait posts on the Web using latent factors . they use features in multiple modalities to characterize the posts and causal inference to eliminate noise .
Outcome: The proposed method can detect clickbait posts on popular social media platforms with good generalization ability.
Causal-Audit: Explicit and Auditable Graph-based Reasoning via Target-Aware Causal Chain Construction (2026.findings-acl)

Copied to clipboard

Challenge: Existing LLM-based methods rely on implicit language-level reasoning, resulting in opaque causal assumptions and fragile predictions.
Approach: They propose an explicit and auditable causal reasoning framework for context-free intervention-based question answering that uses four modular stages rather than implicit end-to-end prediction.
Outcome: The proposed framework outperforms existing LLM-based methods while providing interpretable and auditable causal reasoning traces.
Bayesian Topic Regression for Causal Inference (2021.emnlp-main)

Copied to clipboard

Challenge: a Bayesian topic regression model uses text and numerical information to model outcome variables.
Approach: They propose a Bayesian Topic Regression model that uses both text and numerical information to model an outcome variable.
Outcome: The proposed model recovers ground truth with lower bias than any benchmark model when text and numerical features are correlated.
SLANG: New Concept Comprehension of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Dynamic nature of language limits the adaptability of Large Language Models (LLMs) Traditionally, LLMs are trained on static data, which limits their adaptability .
Approach: They propose a benchmark to integrate novel data and assess LLMs’ ability to comprehend emerging concepts, alongside a causal inference-based approach to enhance LLM comprehension of new phrases and their colloquial context.
Outcome: The proposed model outperforms baseline models in terms of precision and relevance in the comprehension of Internet slang and memes.
Causal Direction of Data Collection Matters: Implications of Causal and Anticausal Learning for NLP (2021.emnlp-main)

Copied to clipboard

Challenge: a meta-analysis of published studies shows that the causal direction of data collection can explain some trends in NLP . semi-supervised learning and domain adaptation performance differ on a number of tasks .
Approach: They argue that the causal direction of the data collection process has nontrivial implications . authors categorize common NLP tasks according to their causal direction . they also empirically assay the validity of the ICM principle for text data .
Outcome: The proposed model can explain differences in semi-supervised learning and domain adaptation performance across settings.
Uncovering Main Causalities for Long-tailed Information Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: Information Extraction (IE) aims to extract structural information from unstructured texts.
Approach: They propose a framework that aims to uncover the main causalities behind data in the view of causal inference.
Outcome: The proposed framework can detect the main causalities behind data in the view of causal inference.
Characterizing and Verifying Scientific Claims: Qualitative Causal Structure is All You Need (2023.emnlp-main)

Copied to clipboard

Challenge: a scientific claim verification requires thorough examination and assessment to ascertain its validity . attention architectures and pre-trained language models fail to establish a comprehensive chain of causal inference .
Approach: They propose a qualitative causal structure-based graph neural network model to facilitate causal reasoning across relevant causally-potent factors.
Outcome: The proposed model outperforms state-of-the-art models by incorporating semantic features . the proposed model is based on a qualitative causal structure .
Can Post-Training Transform LLMs into Causal Reasoners? (2026.findings-acl)

Copied to clipboard

Challenge: Causal inference is a core component of human cognition and requires decision-makers to distinguish between causation and association.
Approach: They propose a dataset comprising seven core causal tasks for training and five diverse test sets and evaluate five different post-training approaches.
Outcome: The proposed model achieves 93.5% accuracy on the CaLM benchmark, compared to 55.4% by OpenAI o3.
Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent Detection (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for multimodal intent detection have two limitations: (i) close entanglement of multimodal semantics with modal structures; (ii) insufficient learning of causal effects of semantic and modality-specific information on the final predictions.
Approach: They propose a Dual-oriented Disentangled Network with Counterfactual Intervention model that decouples semantics-oriented and modality-oriented representations and a Counterfective Intervention Module that applies causal inference to understand causal effects by injecting confounders.
Outcome: The proposed model overcomes key limitations in existing systems by effectively disentangling and utilizing modality-specific and multimodal semantic information.
Context-Aware Reasoning On Parametric Knowledge for Inferring Causal Variables (2025.findings-emnlp)

Copied to clipboard

Challenge: randomized experiments provide strong inferences, but are often infeasible due to ethical or practical constraints.
Approach: They propose a benchmark where the objective is to complete a partial causal graph.
Outcome: The proposed benchmarks show that they can hypothesize backdoor variables between a cause and its effect.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations